Minimum Perfect H Fast N-gram Language

نویسنده

  • Xiao Zhang
چکیده

A new technique is proposed for N-gram language model (LM) retrieval based on minimum perfect hashing (MPH). A hierarchical data structure is used to store N-gram scores in hash tables according to the order of N-grams, and a LM score is retrieved by probing the appropriate hash table slot without collision. Both integer key and character-string key based MPH functions are studied. The proposed MPH-based technique for N-gram LM lookup was evaluated on the Switchboard database and compared with the hierarchical binary search method of ISIP and the combined hash and linear search method of HTK. The proposed MPH-based technique outperformed the ISIP and HTK methods in significantly reduced LM retrieval time for bigram and trigram LMs.

برای دانلود رایگان متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

A fast and memory-efficient N-gram language model lookup method for large vocabulary continuous speech recognition

Recently, minimum perfect hashing (MPH)-based language model (LM) lookup methods have been proposed for fast access of N-gram LM scores in lexical-tree based LVCSR (large vocabulary continuous speech recognition) decoding. Methods of node-based LM cache and LM context pre-computing (LMCP) have also been proposed to combine with MPH for further reduction of LM lookup time. Although these methods...

متن کامل

Exploiting Order - Preserving Speedup N - Gram Language

Minimum Perfect Hashing (MPH) has recently been shown successful in reducing Language Model (LM) lookahead time in LVCSR decoding. In this paper we propose to exploit the orderpreserving (OP) property of a string-key based MPH function to further reduce hashing operation and speed up LM lookahead. A subtree structure is proposed for LM lookahead and an orderpreserving MPH is integrated into the...

متن کامل

Minimal Perfect Hash Rank: Compact Storage of Large N-gram Language Models

In this paper we propose a new method of compactly storing n-gram language models called Minimal Perfect Hash Rank (MPHR) that uses significantly less space than all known approaches. It requires O(n) construction time and allows for O(1) random access of probability values or frequency counts associated with n-grams. We make use of minimal perfect hashing to store fingerprints of n-grams in an...

متن کامل

Randomized Language Models via Perfect Hash Functions

We propose a succinct randomized language model which employs a perfect hash function to encode fingerprints of n-grams and their associated probabilities, backoff weights, or other parameters. The scheme can represent any standard n-gram model and is easily combined with existing model reduction techniques such as entropy-pruning. We demonstrate the space-savings of the scheme via machine tran...

متن کامل

Efficient Minimal Perfect Hash Language Models

The recent availability of large collections of text such as the Google 1T 5-gram corpus (Brants and Franz, 2006) and the Gigaword corpus of newswire (Graff, 2003) have made it possible to build language models that incorporate counts of billions of n-grams. This paper proposes two new methods of efficiently storing large language models that allow O(1) random access and use significantly less ...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

عنوان ژورنال:

دوره   شماره 

صفحات  -

تاریخ انتشار 2002